Papers with inter-annotator agreement

8 papers
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)

Copied to clipboard

Challenge: Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored.
Approach: They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu.
Outcome: The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena .
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)

Copied to clipboard

Challenge: In literature, spoken interactions between characters are of central importance to the narrative.
Approach: They propose to annotate quotations, including their interpersonal structure, for English literary text.
Outcome: The proposed dataset provides a rich view of dialogue structures not available from other available corpora.
TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Existing annotation tools are not efficient for the annotation of corpora and are not error-free.
Approach: They propose to extend existing annotation tools by evaluating their flexibility and efficiency.
Outcome: The proposed system performs platform-independent multimodal annotations and annotates complex textual structures.
PDFAnno: a Web-based Linguistic Annotation Tool for PDF Documents (L18-1)

Copied to clipboard

Challenge: Currently, linguistic annotation tools for PDF documents focus on plain-text documents.
Approach: They propose a web-based linguistic annotation tool for PDF documents . it offers functions for various types of linguistic annotations directly on PDF .
Outcome: The proposed tool can annotate on PDF documents with named entity, dependency relation, and coreference chain.
LMUNIT: Fine-grained Evaluation with Natural Language Unit Tests (2025.findings-emnlp)

Copied to clipboard

Challenge: Using natural language unit tests, language models are costly and noisy, and automated metrics provide only coarse, difficult-to-interpret signals.
Approach: They propose a paradigm that decomposes response quality into explicit, testable criteria and a unified scoring model, LMUnit, which combines multi-objective training across preferences, direct ratings, and natural language rationales.
Outcome: The proposed paradigm significantly improves inter-annotator agreement and enables more effective LLM development workflows.
HighRES: Highlight-based Reference-less Evaluation of Summarization (P19-1)

Copied to clipboard

Challenge: Existing methods for summarizing documents are inconsistent due to the difficulty of manual evaluation.
Approach: They propose a method where summaries are evaluated by multiple annotators against the source document via manually highlighted salient content.
Outcome: The proposed method improves inter-annotator agreement while highlighting differences among systems.
Human vs. Machine Perceptions on Immigration Stereotypes (2024.lrec-main)

Copied to clipboard

Challenge: a growing number of natural language processing models leave aside the language itself . a recent paradigm in the computational linguistics community is training models on specific perspectives of a segment of the population or an individual.
Approach: They propose to use BERT-based classification models to detect stereotypes related to immigrants . they compare models with predictions from GPT-4 and annotated tweets from Spanish Twitter .
Outcome: The proposed models are compared with predictions from the dataset of Spanish Twitter posts containing stereotypes . the models are confident in their predictions and more accurate for implicit stereotypes, the authors show .
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies that analyze unseen domains vary translation systems, annotators, or evaluation conditions, confounding domain effects with human annotation noise.
Approach: They propose to use human error span annotations to evaluate translations of six translation systems across one seen news domain and two unseen technical domains to address these biases.
Outcome: The proposed model improves on the human annotations in two unseen domains and on the news domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations